Back

JMIR Public Health and Surveillance

JMIR Publications Inc.

All preprints, ranked by how well they match JMIR Public Health and Surveillance's content profile, based on 45 papers previously published here. The average preprint has a 0.05% match score for this journal, so anything above that is already an above-average fit. Older preprints may already have been published elsewhere.

1
Creation of a Clinical Decision-Support Tool for Assigning Occupational Disability to United States Air Force Personnel

Uptegraft, C. C.; Witkop, C. T.

2020-05-11 health informatics 10.1101/2020.05.07.20090530 medRxiv
Top 0.1%
13.1%
Show abstract

Occupational dispositions (profiles) are the top reason active duty service members are not medically ready to deploy or fulfill their job responsibilities. An audit across multiple U.S. Air Force (AF) medical treatment facilities revealed significant shortcomings in how medical providers assign profiles. We aimed to create a predictive model and a decision-support tool that estimates profile duration. Using retrospective profiles (n=1,546,805) from the Aeromedical Services Information Management System between 1 Feb 2007 and 31 Jan 2017, we built and validated a decision-support tool that estimates profile length. Multivariate quantile regressions (n=2,575) were performed across five quantiles and six levels of diagnostic specificity for every diagnostic code with 2,100 or more observations. The models universally estimated profile duration with very poor accuracy (pseudoR2 0.000 to 0.168); however, predictive ability was directly correlated with quantile level with minimal variation by diagnostic specificity. Age, O4 to O6+ ranks, very heavy job class, and co-morbid conditions were all significant in more than 25.0% of regressions down all levels of diagnostic specificity. Age, co-morbid conditions, E7-E9 ranks, O4 to O6+ ranks, and light job class all added days to profile duration while E1 to E4 ranks, heavy, and very heavy job class subtracted days. While this study failed to produce an accurate tool, several findings, the indirect correlation between profile duration and very heavy job class and the assignment of durations based on convenient calendar times, warrant further investigation. For now, providers may consult existing decision-support tools when building profiles for AF service members, heeding attention that they were built with non-representative civilian populations. DisclaimerThe views expressed are solely those of the authors and do not reflect the official policy or position of the US Army, US Navy, US Air Force, the Department of Defense, or the US Government.

2
Implementation of a Web-Based Symptom Checker to Manage the Quarantine of the USS Theodore Roosevelt Crew Following a Shipboard Outbreak of SARS-CoV-2

Fenequito, R. L.; Houskamp, D.; Siu, V.; Rorie, J.; Bhatt, N.; Kuang, M.; Luartes, B.; Shell, D.; Reeder, J.; Rainey, K.; Olson, N.

2021-08-28 public and global health 10.1101/2021.08.25.21254738 medRxiv
Top 0.1%
10.1%
Show abstract

IntroductionIn late March 2020, the USS Theodore Roosevelt (TR), a nuclear-powered aircraft carrier, pulled into port in the US territory of Guam to assess the severity of a developing outbreak of COVID-19 aboard the ship. A small staff contingent of 60 personnel from US Naval Hospital (USNH) Guam was tasked with the medical care of 4,079 sailors who were placed in single room quarantine amongst 11 hotels across the island of Guam. With the assistance of the Defense Digital Service, the USNH Guam staff implemented a web-based symptom checker, which allowed for monitoring of developing COVID symptoms, and selective testing of symptomatic individuals. Materials and MethodsSailors from the TR were placed in quarantine or isolation cohorts upon debarking the ship. Sailors not positive for COVID-19 were quarantined amongst 11 hotels on Guam. Sailors positive for COVID-19 were isolated aboard Naval Base Guam (NBG). A retrospective cohort analysis and subgroup analyses were performed on symptom data obtained from sailors in quarantine. The sailors recorded their symptoms and temperature in a web-based symptom checker that assigned a symptom severity score (SSS). Sailors with a SSS >50 were evaluated by a medical provider and re-tested. Data were collected from 4 April 2020 to 1 May 2020. Sailors required two negative tests to exit quarantine and re-embark the ship. The time course, and most common cluster of symptoms associated with a positive COVID-19 PCR test were determined retrospectively after data collection. ResultsThe web-based symptom checker was successful in establishing daily positive contact and symptom monitoring of susceptible individuals in quarantine. 4,079 sailors in quarantine maintained positive contact with medical staff via the symptom checker, with at least 81% of the sailors recording their symptoms on a daily basis. Individuals with high symptom scores were quickly identified and underwent further evaluation and repeat COVID-19 testing. A cohort of 331 sailors tested positive for COVID-19 while in quarantine and recorded symptoms in the symptom checker before and after a positive COVID-19 test. In this cohort, the most frequent symptoms reported prior to a positive test were headache, anosmia, followed by cough. The symptom of anosmia was reported more frequently in sailors positive for COVID-19, compared to a cohort of matched controls. A small medical staff was able to monitor developing symptoms in a large quarantined population, while efficiently allocating resources, preserving personal protective equipment (PPE), and maintaining isolation and social distancing protocols. Conclusions and RelevanceThe application provided a tool for broad health surveillance over a large population while maintaining strict quarantine and social distancing protocols. Highly symptomatic sailors were quickly identified, triaged, and transferred to a higher level of care if indicated. The symptom checker and predictive model generated from the data can be utilized by military and civilian public health officials to triage large populations and make rapid decisions on isolation measures, resource allocation, selective testing.

3
Geographically Weighted Machine Learning for Spatial Prediction of Cancer Prevalence in the United States: A Mixed Method Approach

Sadeghi Naieni Fard, F.; Oppong, J. R.; Tiwari, C.; Boakye, K.; Fard, F.

2026-08-21 public and global health 10.64898/2026.08.18.26360598 medRxiv
Top 0.1%
9.9%
Show abstract

Cancer prevalence is distributed unevenly across regions and caused by the interaction of multiple risk factors. Previous studies focused on the use of global modeling techniques to predict cancer at the county level that overlooks important spatial differences. This study aims to develop geographically weighted machine learning models to predict cancer prevalence at the census tract level in the United States and identify local determinants of cancer burden. First, a scoping review was conducted to find a list of measurable drivers of cancer in the United States. Using this list, the data of these variables for 84415 census tracts were obtained from the Center for Disease Control and Prevention PLACES dataset and other publicly accessible resources. Then, several predictive models, including Ordinary Least Squares (OLS) and Geographically Weighted Regression (GWR), as well as Random Forest, XGBoost, and Deep Neural Network and their geographically weighted counterparts, were developed and compared using the Coefficient of Determination, Root Mean Square Error, and Absolute Error. Results presented that geographically weighted models outperformed other methods, and geographically weighted XGBoost achieved the strongest and most consistent overall performance with pseudo-R2 ranging between 0.89 and 0.98. Feature importance analysis of this model illustrated that most important cancer drivers changed location by location. Aged people, racial composition, preventative behaviors, and metabolic conditions such as diabetes, hypertension, and high cholesterol were determined as influential predictors, although their relative importance varied across regions. These findings revealed the value of localized models at a small geographic scale to identify regional cancer risk patterns and help the allocation of proper resources to hotspot areas. Keywords: Cancer prevalence, Census tracts, geographically weighted machine learning models, Deep neural network, XGBoost, Random Forest, Ordinary Least Squares, risk factor, determinant

4
Smartphone-based App to Assess Diabetic Peripheral Neuropathy

Adenekan, R. A. G.; Adenekan, A. E.; Leung, K. K.; Muppidi, S.; Sakamuri, S.; Tan, M.; Tsai, S. A.; Osikomaiya, M.; Okamura, A. M.; Nunez, C. M.; Kim, S. H.; Yoshida, K. T.

2025-09-02 endocrinology 10.1101/2025.08.28.25333808 medRxiv
Top 0.1%
9.9%
Show abstract

BackgroundDiabetic peripheral neuropathy (DPN) affects approximately 50% of individuals with diabetes and is a risk factor for amputations. Unfortunately, foot exams and screening tools are inconsistent and miss early-stage nerve damage. A smartphone-based application that delivers controlled vibrations, records patient responses, and computes a vibration perception thresh-old (SVPT) may present an accessible, precise monitoring avenue. This study assesses the clinical relevance and precision of SVPTs for measuring large-fiber sensory deficits in patients with diabetes. MethodsWe measured SVPTs in 71 patients with pre-diabetes or diabetes and compared their efficacy with tuning fork exams. We analyzed the correlation between SVPT and Rydel-Seiffer tuning fork (RSTF) scores, along with their relationship with clinical DPN markers such as hemoglobin A1c (HbA1c), age, and disease duration using multivariable linear regression. ResultsSVPTs moderately correlated with RSTF scores (Rs = -0.43, p = 0.0019). Among adults aged 50 to 69, SVPTs correlated significantly with clinical markers (F (4, 29) = 4.76, p = 0.00447, Multiple R2 = 0.396, Adjusted R2 = 0.313,{varepsilon} = 0.167). The interaction between age and HbA1c was positively associated with SVPTs ({beta} = 0.118, p = 0.001), while SVPTs were negatively associated with diabetes duration ({beta} = -0.098, p = 0.003). ConclusionsWe present a clinically relevant, patient-operated smartphone application for large-fiber sensory monitoring, tested on patients with varying DPN risk. This novel platform has the potential to provide a precise, reliable, and accessible avenue for identifying individuals at risk of developing DPN complications, prior to overt clinical manifestation.

5
Modeling the progression of SARS-CoV-2 infection in patients with COVID-19 risk factors through predictive analysis

Leon-Abarca, J. A.; Musayon, E. I.

2020-09-12 public and global health 10.1101/2020.07.14.20154021 medRxiv
Top 0.1%
9.9%
Show abstract

With almost a third of adults being obese, another third hypertense and almost a tenth affected by diabetes, Latin American countries could see an elevated number of severe COVID-19 outcomes. We used the Open Dataset of Mexican patients with COVID-19 suspicion who had a definite RT-PCR result to develop a statistical model that evaluated the progression of SARS-CoV-2 infection in the population. We included patients of all ages with every risk factor provided by the dataset: asthma, chronic obstructive pulmonary disease, smoking, diabetes, obesity, hypertension, immunodeficiencies, chronic kidney disease, cardiovascular diseases, and pregnancy. The dataset also included an unspecified category for other risk factors that were not specified as a single variable. To avoid excluding potential patients at risk, that category was included in our analysis. Due to the nature of the dataset, the calculation of a standardized comorbidity index was not possible. Therefore, we treated risk factors as a categorical variable with two categories: absence of risk factors and the presence of at least one risk factor in accordance with previous epidemiological reports. Multiple logistic regressions were carried out to associate sex, risk factors, and age as a continuous variable (and the interaction that accounted for increasing diseases with older ages); and SARS-CoV-2 infection as the dependent zero-one binomial variable. Post estimation predictive marginal analysis was performed to generate probability trends along 95% confidence bands. This analysis was repeated several times through the course of the pandemic since the first record provided in their repository (April 12, 2020) to one month after the end of the state of sanitary emergency (the last date analyzed: June 27, 2020). After processing, the last measurement included 464,389 patients. The baseline analysis on April 12 revealed that people 35 years and older with at least one risk factor had a lower risk of SARS-CoV-2 infection in comparison to patients without risk factors (Figure 1). One month before the end of the nationwide state of emergency this age threshold was found at 50 years (May 2, 2020) and it shifted to 65 years on May 30. Two weeks after the end of the public emergency (June 13, 2020) the trends converged at 80 years and one week later (June 27, 2020) every male and female patient with at least one risk factor had a higher risk of SARS-CoV-2 infection compared to people without risk factors. Through the course of the COVID-19 pandemic, all four probability curves shifted upwards as a result of progressive disease spread. O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=140 SRC="FIGDIR/small/20154021v1_fig1.gif" ALT="Figure 1"> View larger version (34K): org.highwire.dtl.DTLVardef@a5bca3org.highwire.dtl.DTLVardef@1039176org.highwire.dtl.DTLVardef@14314f7org.highwire.dtl.DTLVardef@1159d32_HPS_FORMAT_FIGEXP M_FIG O_FLOATNOImage 1C_FLOATNO SARS-CoV-2 infection probability progression by date. Green and orange trends: Patients with at least one risk factor. Gray trends: Patients with no risk factors. C_FIG In conclusion, we found our model could monitor accurately the probability of SARS-CoV-2 infection in relation to age, sex, and the presence of at least one risk factor. Also, because the model can be applied to any particular political region within Mexico, it could help evaluate the contagion spread in specific vulnerable populations. Further studies are needed to determine the underlying nature of the mechanisms behind such observations.

6
Predicting COVID-19 case status from self-reported symptoms and behaviors using data from a massive online survey

Srivastava, M.; Reinhart, A.; Mejia, R.

2023-02-07 public and global health 10.1101/2023.02.03.23285405 medRxiv
Top 0.1%
9.7%
Show abstract

AO_SCPLOWBSTRACTC_SCPLOWWith the varying availability of RT-PCR testing for COVID-19 across time and location, there is a need for alternative methods of predicting COVID-19 case status. In this study, multiple machine learning (ML) models were trained and assessed for their ability to accurately predict the COVID-19 case status using US COVID-19 Trends and Impact Survey (CTIS) data. The CTIS includes information on testing, symptoms, demographics, behaviors, and vaccination status. The best performing model was XGBoost, which achieved an F1 score of{approx} 94% in predicting whether an individual was COVID-19 positive or negative. This is a notable improvement on existing models for predicting COVID-19 case status and demonstrates the potential for ML methods to provide policy-relevant estimates.

7
Enabling a Learning Public Health System: Enhanced Surveillance of HIV and Other Sexually Transmitted Infections

Reyes Nieva, H.; Zucker, J.; Tucker, E.; Castor, D.; Yin, M. T.; Gordon, P.; Elhadad, N.

2024-04-12 health informatics 10.1101/2024.04.10.24305612 medRxiv
Top 0.1%
9.6%
Show abstract

Sexually transmitted infections (STIs) continue to pose a substantial public health challenge in the United States (US). Surveillance, a cornerstone of disease control and prevention, can be strengthened to promote more timely, efficient, and equitable practices by incorporating health information exchange (HIE) and other large-scale health data sources into reporting. New York City patient-level electronic health record data between January 1, 2018 and June 30, 2023 were obtained from Healthix, the largest US public HIE. Healthix data were linked to neighborhood-level information from the American Community Survey. In this cross-sectional study, we compared patients who received a test or tested positive for chlamydia, gonorrhea, and/or HIV with patients who were untested or tested negative, respectively, using generalized estimating equations with logit function and robust standard errors. Among 1,519,121 tests performed for chlamydia, 1,574,772 for gonorrhea, and 1,200,560 for HIV, 2%, 0.6% and 0.3% were positive for chlamydia, gonorrhea, and HIV, respectively. Chlamydia and gonorrhea co-occurred in 1,854 cases (7% of chlamydia and 21% of gonorrhea total cases). Testing behavior was often incongruent with geographic and sociodemographic patterns of positive cases. For example, people living in areas with the highest levels of poverty were less likely to test for gonorrhea but almost twice as likely to test positive compared to those in low poverty areas. Regional HIE enabled review of testing and cases using granular and complementary data not typically available given existing reporting practices. Enhanced surveillance spotlights potential incongruencies between testing patterns and STI risk in certain populations, signaling potential under- and over-testing. These and future insights derived from HIE data may be used to continuously inform public health practice and drive further improvements in provision and evaluation of services and programs.

8
Using Over-the-Counter Retail Medication Sales to Detect and Track Influenza-Like Illnesses Including Novel Diseases

Ye, Y.; Espino, J.; Aronis, J. M.; Hochheiser, H.; Michaels, M. G.; Cooper, G. F.

2025-11-09 public and global health 10.1101/2025.11.07.25339792 medRxiv
Top 0.1%
9.1%
Show abstract

Timely detection of emerging disease outbreaks is critical for effective public health response. Traditional surveillance systems often rely on clinical or laboratory-confirmed data, which may delay early detection. In this study, we evaluate the potential of the National Retail Data Monitor (NRDM), which tracks over-the-counter (OTC) health-related product sales, as a tool for public health surveillance. Using data from Allegheny County, Pennsylvania (2016-2021), we developed a probabilistic model that estimates daily influenza-like illness (ILI) activity based on purchasing patterns. To train the model, we applied an expectation-maximization (EM) algorithm to estimate the conditional probabilities of OTC product purchases given ILI or not ILI, which forms the basis of our model. To test the model, we used that model and daily OTC sales to estimate the number of purchases each day due to ILI. The estimated ILI counts showed moderately strong correlation (r = 0.66) with ILI counts derived from emergency departments (EDs) in the same region. We also monitor the extent to which our model predicts the data well; if the current prediction is poor, relative to historical predictions, it raises the prospect that there is an outbreak of an unusual disease in the population, perhaps even a novel disease. Our findings suggest that surveillance of OTC products can effectively identify unusual health-related activity and provide early insights of potential known and novel outbreaks through changes in consumer purchasing behavior.

9
Factors that influenced testing positive and dying from COVID-19 in the West South-Central Division of the United States and its Effects on Future Public Health Policy

Kaliba, A. R.

2025-08-06 public and global health 10.1101/2025.08.01.25332736 medRxiv
Top 0.1%
9.0%
Show abstract

ObjectivesExamine the relationship between various demographic characteristics and meso variables measuring social vulnerability, religiosity, political partisanship, and the built environment on the probability of testing positive for COVID-19 and dying after testing positive for the virus. MethodsThe individual-level variables are from the data collected by the Louisiana Department of Health and Hospitals. Meso variables were sourced from various platforms. The data are analyzed using a spatial bivariate probit model with copulas in a generalized additive model framework. ResultsThe main results suggest a strong and positive association between individual-level covariates and testing positive for the virus and dying from it. The effects of social vulnerability, religiosity, political partisanship, and the built environment varied non-linearly; their effects were within a given critical range. ConclusionsTo mitigate the impact of future pandemics like COVID-19, public health policies should focus on addressing existing health disparities, fostering meaningful engagement with community institutions and diverse leaders, and applying proven and scientific public health considerations, while minimizing the influence of political ideologies and culture.

10
Estimated prevalence of multiple chronic conditions throughout adulthood using data from the All of Us Research Program

Li, X.; Dreisbach, C.; Gustafson, C. M.; Murali, K. P.; Koleck, T. A.

2024-10-18 primary care research 10.1101/2024.10.17.24315661 medRxiv
Top 0.1%
9.0%
Show abstract

Estimation of multiple chronic condition (MCC) prevalence throughout adulthood provides a critical reflection of MCC burden. We analyzed electronic health record codes for 58 conditions to estimate MCC prevalence for All of Us (AoU) Research Program adult participants (N=242,828). Approximately 76% of AoU participants were diagnosed with MCCs, with over 40% having 6 or more conditions and prevalence increasing with age; the most frequently occurring MCC combinations varied by age category (i.e., mental health conditions in early adulthood and physical health conditions in middle adulthood through advanced old age). We report notable prevalence of MCC throughout adulthood and variability in MCC condition combinations by age category in AoU participants. These findings highlight the need for targeted, innovative care modalities and population health initiatives to address MCC burden throughout adulthood.

11
A Snapshot of COVID-19 Incidence, Hospitalizations, and Mortality from Indirect Survey Data in China in January 2023

Ramirez, J. M.; Diaz-Aranda, S.; Aguilar, J.; Ojo, O.; Lillo, R. E.; Fernandez Anta, A.

2023-02-26 public and global health 10.1101/2023.02.22.23286167 medRxiv
Top 0.1%
8.9%
Show abstract

In this work we estimate the incidence of COVID-19 in China using online indirect surveys (which preserve the privacy of the participants). The indirect surveys deployed collect data on the incidence of COVID-19, asking the participants about the number of cases, deaths, vaccinated, and hospitalized that they know. The incidence of COVID-19 (cases, deaths, etc.) is then estimated using a modified Network Scale-up Method (NSUM). Survey responses (100, 200 and 1,000, respectively) were collected from Australia, the UK, and China in January 2023. The estimates in Australia and the UK are compared with official data, showing that they are in the confidence intervals or rather close. Cronbachs alpha values also indicate good confidence in the estimates. The estimates obtained in China are, among others, that 91% of the population is vaccinated, almost 80% had been infected in the last month, and almost 3% in the last 24 hours.

12
Interim evaluation of Google AI forecasting for COVID-19 compared with statistical forecasting by human intelligence in the first week

Kurita, J.; Sugawara, T.; Ohkusa, Y.

2020-12-18 public and global health 10.1101/2020.12.16.20248358 medRxiv
Top 0.1%
8.3%
Show abstract

BackgroundSince June, Google (Alphabet Inc.) has provided forecasting for COVID-19 outbreak by artificial intelligence (AI) in the USA. In Japan, they provided similar services from November, 2020. ObjectWe compared Google AI forecasting with a statistical model by human intelligence. MethodWe regressed the number of patients whose onset date was day t on the number of patients whose past onset date was 14 days prior, with information about traditional surveillance data for common pediatric infectious diseases including influenza, and prescription surveillance 7 days prior. We predicted the number of onset patients for 7 days, prospectively. Finally, we compared the result with Googles AI-produced forecast. We used the discrepancy rate to evaluate the precision of prediction: the sum of absolute differences between data and prediction divided by the aggregate of data. ResultsWe found Google prediction significantly negative correlated with the actual observed data, but our model slightly correlated but not significant. Moreover, discrepancy rate of Google prediction was 27.7% for the first week. The discrepancy rate of our model was only 3.47%. Discussion and ConclusionResults show Googles prediction has negatively correlated and greater difference with the data than our results. Nevertheless, it is noteworthy that this result is tentative: the epidemic curve showing newly onset patients was not fixed.

13
Preventive dental visits and health literacy in patients with diabetes: a nationwide cross-sectional study

Saito, K.; Kawai, Y.; Ishikawa, H.; Tabuchi, T.; Kuwahara, K.

2024-07-03 endocrinology 10.1101/2024.07.01.24309770 medRxiv
Top 0.1%
8.2%
Show abstract

AimThis cross-sectional study examined the association between health literacy and preventive dental visits in patients with diabetes. MethodsWe used cross-sectional data from the Japan COVID-19 and Society Internet Survey (JACSIS), a web-based nationwide survey. The participants were 1,441 patients reporting to have diabetes in 2020. Health literacy was measured using the validated scales for health literacy. Preventive dental visits in the past 12 months were self-reported. We estimated the multivariable-adjusted prevalence ratio (PR) for preventive dental visits. ResultsOver 50% of the patients had preventive dental visits in the past 12 months, and approximately one-third had high health literacy. Compared with the low health literacy group, the high health literacy group was more likely to engage in preventive dental visits (the multivariable-adjusted PR associated with high health literacy: 1.12 [95% confidence interval: 1.01 to 1.23]). Similar results were obtained when health literacy was treated as a continuous variable. ConclusionsThe present data from the JACSIS showed that health literacy was positively associated with preventive dental visits among patients with diabetes.

14
External validation of Finnish Diabetes Risk Score (FINDRISC) and Latin American FINDRISC for screening of undiagnosed dysglycemia: analysis in a Peruvian hospital health care workers sample.

Yovera-Aldana, M.; Mezones-Holguin, E.; Agüero-Zamora, R.; Damas-Casani, L.; Uriol-Llanos, B.; Espinoza-Morales, F.; Soto-Becerra, P.; Ticse-Aguirre, R.

2024-02-18 endocrinology 10.1101/2024.02.16.24302929 medRxiv
Top 0.1%
8.1%
Show abstract

AimsTo evaluate the external validity of Finnish diabetes risk score (FINDRISC) and Latin American FINDRISC (LAFINDRISC) for undiagnosed dysglycemia in hospital health care workers. MethodsWe carried out a cross-sectional study on health workers without a prior history of diabetes mellitus (DM). Undiagnosed dysglycemia (prediabetes or diabetes mellitus) was defined using fasting glucose and two-hour oral glucose tolerance test. LAFINDRISC is an adapted version of FINDRISC with different waist circumference cut-off points. We calculated the area under the receptor operational characteristic curve (AUROC) and explored the best cut-off point. ResultsWe included 549 participants in the analysis. The frequency of undiagnosed dysglycemia was 17.8%. The AUROC of LAFINDRISC and FINDRISC were 71.5% and 69.2%; p=0.007, respectively. The optimal cut-off for undiagnosed dysglycemiaaccording to Index Youden was [≥] 11 in LAFINDRISC (Sensitivity: 78.6%; Specificity: 51.7%) and [≥]12 in FINDRISC (Sensitivity: 70.4%; Specificity: 53.9%) ConclusionThe discriminative capacity of both questionnaires is good for the diagnosis of dysglycemia in the healthcare personnel of the Maria Auxiliadora hospital. The LAFINDRISC presented a small statistical difference, nontheless clinically similar, since there was no difference by age or sex. Further studies in the general population are required to validate these results.

15
Hybrid Fuzzy Logic and Logistic Regression Model with Recursive Feature Elimination for Enhanced Prediction and Clinical Decision Support in Type 2 Diabetes Mellitus Among Adults Aged 35 to 45

Dadashkarimi, M.

2025-07-29 endocrinology 10.1101/2025.07.29.25332372 medRxiv
Top 0.1%
8.0%
Show abstract

Withdrawal StatementThe authors have withdrawn their manuscript owing to errors in the experimental design that affect the integrity of the results. Therefore, the authors do not wish this work to be cited as reference for the project. If you have any questions, please contact the corresponding author.

16
Prediction of COVID-19 infection risk using personal mobile location data only

Jang, A.; Kim, S.; Baek, H.; Kim, H.; Park, H.-L.

2023-08-24 infectious diseases 10.1101/2023.08.22.23294419 medRxiv
Top 0.1%
8.0%
Show abstract

Predicting an individuals risk of infectious disease is a critical technology in infectious disease response. During the COVID-19 pandemic, identifying and isolating individuals at high risk of infection was an essential task for epidemic control. We introduce a new machine learning model that predicts the risk of COVID-19 infection using only individuals mobile cell tower location information. This model distinguishes the cell tower location information of an individual into residential and non-residential areas and calculates whether the cell tower locations overlapped with other individuals. It then generates various variables from the information of overlapping and predicts the possibility of COVID-19 infection using a machine learning algorithm. The predictive model we developed showed performance comparable to models using individuals clinical information. This predictive model, which can be used to predict infections of diseases with asymptomatic infections such as COVID-19, has the advantage of supplementing the limitations of existing infectious disease prediction models that use symptoms and other information.

17
Accurate Covid-19 prevalence measurement in the field

Georganas, S.; Velias, A.; Vandoros, S.

2021-01-11 public and global health 10.1101/2021.01.10.21249298 medRxiv
Top 0.1%
8.0%
Show abstract

Timely, accurate epidemic figures are necessary for informed policy. In the Covid-19 pandemic, mismeasurement can lead to tremendous waste, in health or economic output. "Random" testing is commonly used to estimate virus prevalence, reporting daily positivity rates. However, since testing is necessarily voluntary, all "random" tests done in the field suffer from selection bias. This bias, unlike standard polling biases, goes beyond demographical representativeness and cannot be corrected by oversampling (i.e. selecting people without symptoms to test). Using controlled, incentivized experiments on a sample of all ages, we show that people who feel symptoms are up to 33 times more likely to seek testing. The bias in testing propensities leads to sizable prevalence bias: test positivity is inflated by up to five times, even if testing is costless. This effect varies greatly across time and age groups, making comparisons over time and across countries misleading. We validate our results using the REACT study in the UK and find that positivity figures have indeed a very large and time varying bias. We present calculations to debias positivity rates, but importantly, suggest a parsimonious way to sample the population bypassing the bias altogether. Our estimation is both real-time and consistently close to true values. These results are relevant for all epidemics, besides covid-19, when carriers have informative beliefs about their own status.

18
Forecasting COVID-19 cases in US states using reconstructed incidence data

Nash, R. K.; Bhatia, S.; Wardle, J.; Cori, A.; Nouvellet, P.

2025-10-21 public and global health 10.1101/2025.10.20.25338346 medRxiv
Top 0.1%
8.0%
Show abstract

Branching process models are commonly used in infectious disease forecasting and often rely on daily incidence data, but their utility can be restricted if incidence is not reported daily or if reporting becomes less frequent during prolonged outbreaks. In this study, an Expectation Maximisation algorithm is used to reconstruct the daily incidence of COVID-19 cases from weekly case counts. Using data from 13 US states that maintained mostly daily reporting of COVID-19 cases from March 2020 to February 2022, we evaluate forecasting performance by comparing models using the true daily incidence with those using reconstructed daily incidence. Our results show that forecasts generated from reconstructed incidence perform equally well as those generated using true daily incidence. These findings demonstrate the viability of using reconstructed incidence data for real-time forecasting, which could be particularly useful in scenarios where maintaining daily reporting is unsustainable or in settings with limited surveillance capacity.

19
Exploring the Temporal Dynamics of County-Level Vulnerability Factors on COVID-19 Outcomes

Zhang, J.; Choi, D.; Patel, S. A.; Ho, J. C.

2021-11-26 public and global health 10.1101/2021.11.24.21266757 medRxiv
Top 0.1%
7.9%
Show abstract

As the outbreak of COVID-19 has become a severe worldwide pandemic, every country fights against the spread of this deadly disease with incredible efforts. There are numerous researches along with every conceivable dimension for COVID-19. Among these researches, different demographic and contextual factors of populations and communities also play an essential role in providing more information for decision-makers. This paper mainly utilizes existing data on county contextual factors at the United States county-level to develop a model that can capture the dynamic trajectory of COVID-19 (i.e., cases) and its impacts across the United States. Moreover, our methods applied to contextual data achieves better results compared with existing measures of vulnerability.

20
Trends and Disparities in Newer GLP1 Receptor Agonist Initiation among Real-World Adult Patients Eligible for Obesity Treatment

Radwan, R. M.; Lee, Y. A.; Kotecha, P.; Write, D.; Hernandez, I.; Ramon, R.; Donahoo, W. T.; Chen, Y.; Ellen, J.; Jiang, B.; Guo, J.

2025-01-20 endocrinology 10.1101/2025.01.20.25320839 medRxiv
Top 0.1%
7.9%
Show abstract

AimsTo characterize trends in the initiation of newer anti-obesity medications (AOMs) and determine factors associated with their use among obese/overweight populations. Materials and methodsThis retrospective study utilized electronic health record data from OneFlorida+ (2015-2024). Adults eligible for AOMs were included, defined as having a BMI [≥]30 kg/m{superscript 2} or a BMI of 27-29.9 kg/m{superscript 2} with at least one obesity-related comorbidity. The primary outcome was the initiation of newer AOMs, specifically glucagon-like peptide-1 receptor agonists (GLP-1 RAs) including liraglutide, semaglutide, and tirzepatide. Trends across years were examined, and a multivariable logistic regression identified sociodemographic, clinical, and healthcare utilization factors associated with AOM initiation. ResultsOf 319,949 adults, 1.8% initiated newer AOMs. Semaglutide accounted for 77.9% of initiations, tirzepatide 19.7%, and liraglutide 17.8%. Initiation trends showed liraglutide uptake peaked at 5% in 2018 but declined afterward, while semaglutide and tirzepatide uptake increased exponentially since 2022. Odds of initiation were lower for Black (aOR (95% CI): 0.87 [0.80- 0.94]) and Hispanic (0.84 [0.78-0.91]) groups vs. Whites, and for Medicaid (0.69 [0.63-0.76]) and uninsured (0.81 [0.74-0.87]) patients vs. privately insured. Higher odds were associated with being female, middle-aged, having more outpatient visits, and visiting endocrinologists. ConclusionsThe initiation of newer AOMs among overweight and obese populations remains low, but uptake has increased exponentially since 2022. Our findings reveal significant disparities in obesity care, highlighting the importance of addressing inequities in AOM access to improve obesity outcomes.